A conclusive methodology for rating OCR performance
Identifieur interne : 001348 ( Main/Exploration ); précédent : 001347; suivant : 001349A conclusive methodology for rating OCR performance
Auteurs : Nathan E. Brener [États-Unis] ; S. S. Iyengar [États-Unis] ; O. S. Pianykh [États-Unis]Source :
- Journal of the American Society for Information Science and Technology [ 1532-2882 ] ; 2005-10.
Descripteurs français
- Pascal (Inist)
English descriptors
- KwdEn :
Abstract
One of the most challenging topics in the automatic document rating process is the development of a rating scheme for the image quality of documents. As part of the Department of Energy (DOE) document declassification program, we have developed a generalized rating system to predict the optical character recognition (OCR) accuracy level that is achieved when processing a document. The need for such a system emerged from the declassification of degraded, typewriter‐era documents, which is currently a time‐consuming manual process. This article presents the statistical analysis of the most influential document quality features affecting OCR accuracy, develops consistent predictive models for four currently used OCR engines, and studies the applicability of different OCR products to the DOE document declassification process. This study is expected to lead to an efficient and completely automated document declassification system.
Url:
DOI: 10.1002/asi.20214
Affiliations:
Links toward previous steps (curation, corpus...)
- to stream Istex, to step Corpus: 000000
- to stream Istex, to step Curation: 000000
- to stream Istex, to step Checkpoint: 000C57
- to stream Main, to step Merge: 001384
- to stream PascalFrancis, to step Corpus: 000391
- to stream PascalFrancis, to step Curation: 000395
- to stream PascalFrancis, to step Checkpoint: 000444
- to stream Main, to step Merge: 001474
- to stream Main, to step Curation: 001348
Le document en format XML
<record><TEI wicri:istexFullTextTei="biblStruct"><teiHeader><fileDesc><titleStmt><title xml:lang="en">A conclusive methodology for rating OCR performance</title>
<author><name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
</author>
<author><name sortKey="Iyengar, S S" sort="Iyengar, S S" uniqKey="Iyengar S" first="S. S." last="Iyengar">S. S. Iyengar</name>
</author>
<author><name sortKey="Pianykh, O S" sort="Pianykh, O S" uniqKey="Pianykh O" first="O. S." last="Pianykh">O. S. Pianykh</name>
</author>
</titleStmt>
<publicationStmt><idno type="wicri:source">ISTEX</idno>
<idno type="RBID">ISTEX:8B4039C936354C5EA2A1027A0E8248A4629FAC7B</idno>
<date when="2005" year="2005">2005</date>
<idno type="doi">10.1002/asi.20214</idno>
<idno type="url">https://api.istex.fr/document/8B4039C936354C5EA2A1027A0E8248A4629FAC7B/fulltext/pdf</idno>
<idno type="wicri:Area/Istex/Corpus">000000</idno>
<idno type="wicri:Area/Istex/Curation">000000</idno>
<idno type="wicri:Area/Istex/Checkpoint">000C57</idno>
<idno type="wicri:doubleKey">1532-2882:2005:Brener N:a:conclusive:methodology</idno>
<idno type="wicri:Area/Main/Merge">001384</idno>
<idno type="wicri:source">INIST</idno>
<idno type="RBID">Pascal:06-0252013</idno>
<idno type="wicri:Area/PascalFrancis/Corpus">000391</idno>
<idno type="wicri:Area/PascalFrancis/Curation">000395</idno>
<idno type="wicri:Area/PascalFrancis/Checkpoint">000444</idno>
<idno type="wicri:doubleKey">1532-2882:2005:Brener N:a:conclusive:methodology</idno>
<idno type="wicri:Area/Main/Merge">001474</idno>
<idno type="wicri:Area/Main/Curation">001348</idno>
<idno type="wicri:Area/Main/Exploration">001348</idno>
</publicationStmt>
<sourceDesc><biblStruct><analytic><title level="a" type="main" xml:lang="en">A conclusive methodology for rating OCR performance</title>
<author><name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
<affiliation wicri:level="2"><country xml:lang="fr">États-Unis</country>
<placeName><region type="state">Louisiane</region>
</placeName>
<wicri:cityArea>Department of Computer Science, Louisiana State University, Baton Rouge</wicri:cityArea>
</affiliation>
<affiliation wicri:level="1"><country wicri:rule="url">États-Unis</country>
</affiliation>
</author>
<author><name sortKey="Iyengar, S S" sort="Iyengar, S S" uniqKey="Iyengar S" first="S. S." last="Iyengar">S. S. Iyengar</name>
<affiliation wicri:level="2"><country xml:lang="fr">États-Unis</country>
<placeName><region type="state">Louisiane</region>
</placeName>
<wicri:cityArea>Department of Computer Science, Louisiana State University, Baton Rouge</wicri:cityArea>
</affiliation>
</author>
<author><name sortKey="Pianykh, O S" sort="Pianykh, O S" uniqKey="Pianykh O" first="O. S." last="Pianykh">O. S. Pianykh</name>
<affiliation wicri:level="2"><country xml:lang="fr">États-Unis</country>
<placeName><region type="state">Louisiane</region>
</placeName>
<wicri:cityArea>Department of Computer Science, Louisiana State University, Baton Rouge</wicri:cityArea>
</affiliation>
</author>
</analytic>
<monogr></monogr>
<series><title level="j">Journal of the American Society for Information Science and Technology</title>
<title level="j" type="abbrev">J. Am. Soc. Inf. Sci.</title>
<idno type="ISSN">1532-2882</idno>
<idno type="eISSN">1532-2890</idno>
<imprint><publisher>Wiley Subscription Services, Inc., A Wiley Company</publisher>
<pubPlace>Hoboken</pubPlace>
<date type="published" when="2005-10">2005-10</date>
<biblScope unit="volume">56</biblScope>
<biblScope unit="issue">12</biblScope>
<biblScope unit="page" from="1274">1274</biblScope>
<biblScope unit="page" to="1287">1287</biblScope>
</imprint>
<idno type="ISSN">1532-2882</idno>
</series>
<idno type="istex">8B4039C936354C5EA2A1027A0E8248A4629FAC7B</idno>
<idno type="DOI">10.1002/asi.20214</idno>
<idno type="ArticleID">ASI20214</idno>
</biblStruct>
</sourceDesc>
<seriesStmt><idno type="ISSN">1532-2882</idno>
</seriesStmt>
</fileDesc>
<profileDesc><textClass><keywords scheme="KwdEn" xml:lang="en"><term>Automatic processing</term>
<term>Document processing</term>
<term>Image evaluation</term>
<term>Image processing</term>
<term>Image quality</term>
<term>Optical character recognition</term>
<term>Parameter estimation</term>
<term>Regression analysis</term>
<term>System performance</term>
</keywords>
<keywords scheme="Pascal" xml:lang="fr"><term>Analyse régression</term>
<term>Estimation paramètre</term>
<term>Evaluation image</term>
<term>Performance système</term>
<term>Qualité image</term>
<term>Reconnaissance optique caractère</term>
<term>Traitement automatique</term>
<term>Traitement document</term>
<term>Traitement image</term>
</keywords>
</textClass>
<langUsage><language ident="en">en</language>
</langUsage>
</profileDesc>
</teiHeader>
<front><div type="abstract" xml:lang="en">One of the most challenging topics in the automatic document rating process is the development of a rating scheme for the image quality of documents. As part of the Department of Energy (DOE) document declassification program, we have developed a generalized rating system to predict the optical character recognition (OCR) accuracy level that is achieved when processing a document. The need for such a system emerged from the declassification of degraded, typewriter‐era documents, which is currently a time‐consuming manual process. This article presents the statistical analysis of the most influential document quality features affecting OCR accuracy, develops consistent predictive models for four currently used OCR engines, and studies the applicability of different OCR products to the DOE document declassification process. This study is expected to lead to an efficient and completely automated document declassification system.</div>
</front>
</TEI>
<affiliations><list><country><li>États-Unis</li>
</country>
<region><li>Louisiane</li>
</region>
</list>
<tree><country name="États-Unis"><region name="Louisiane"><name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
</region>
<name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
<name sortKey="Iyengar, S S" sort="Iyengar, S S" uniqKey="Iyengar S" first="S. S." last="Iyengar">S. S. Iyengar</name>
<name sortKey="Pianykh, O S" sort="Pianykh, O S" uniqKey="Pianykh O" first="O. S." last="Pianykh">O. S. Pianykh</name>
</country>
</tree>
</affiliations>
</record>
Pour manipuler ce document sous Unix (Dilib)
EXPLOR_STEP=$WICRI_ROOT/Ticri/CIDE/explor/OcrV1/Data/Main/Exploration
HfdSelect -h $EXPLOR_STEP/biblio.hfd -nk 001348 | SxmlIndent | more
Ou
HfdSelect -h $EXPLOR_AREA/Data/Main/Exploration/biblio.hfd -nk 001348 | SxmlIndent | more
Pour mettre un lien sur cette page dans le réseau Wicri
{{Explor lien |wiki= Ticri/CIDE |area= OcrV1 |flux= Main |étape= Exploration |type= RBID |clé= ISTEX:8B4039C936354C5EA2A1027A0E8248A4629FAC7B |texte= A conclusive methodology for rating OCR performance }}
This area was generated with Dilib version V0.6.32. |